Skip to content

Move the QUIC transport to ngtcp2 and add H3EarlyData for 0-RTT - #12

Closed
xl32 wants to merge 5 commits into
machine-moon:trunkfrom
xl32:ngtcp2-0rtt
Closed

xl32 wants to merge 5 commits into
machine-moon:trunkfrom
xl32:ngtcp2-0rtt

Conversation

@xl32

@xl32 xl32 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

OpenSSL's own QUIC server drops 0-RTT packets, so 0-RTT cannot work on
it. This change runs the QUIC transport on ngtcp2 and keeps OpenSSL only
for the TLS 1.3 handshake, through its QUIC TLS API. This is the same
split that nginx uses.

The h3q_ API does not change. Its implementation now reads the UDP
socket, maps connection IDs to connections, answers Retry and Version
Negotiation, and runs the ngtcp2 timers on a monotonic clock. The BIO
filter that recovered peer addresses from the OpenSSL listener is
removed. The engine does not copy stream data: it reports each
acknowledgement through stream_acked, and nghttp3 releases the buffer
when the peer acknowledges it.

Connection close follows RFC 9000 section 10. A local close keeps its
CONNECTION_CLOSE packet and resends it, at most once per PTO, for three
PTOs. An idle timeout closes silently. The transport idle timer runs
2 seconds past H3IdleTimeout, so the module's clean close comes first.

flush_nghttp3 now blocks a stream again each time a write blocks.
nghttp3_conn_resume_stream reschedules a blocked stream, and the old
code then offered the same stream forever. OpenSSL's large write buffer
hid this; ngtcp2 reports the real peer flow control.

H3EarlyData (default off) turns on 0-RTT. With it on, a returning client
sends its first request with the handshake and gets the response one
round trip sooner. Only safe methods run before the handshake completes;
other methods wait for it, because 0-RTT data can be replayed (RFC 8470).
The TLS layer sets SSL_OP_NO_ANTI_REPLAY, because OpenSSL's own
anti-replay is built for TLS over TCP and refuses QUIC resumption with
early data.

xl32 added 3 commits September 8, 2026 11:39
A host with "h3" in Protocols now serves HTTP/3 with the certificate
mod_ssl already resolved for it (SSLCertificateFile and mod_md alike),
the way mod_http2 rides mod_ssl. The certificate and key are loaded in
the ap_ssl_add_cert_files hook, during startup while httpd still runs
privileged, so a key readable only by root loads for QUIC as it does for
mod_ssl instead of failing in the unprivileged child.

H3CertificatePath and H3CertificateKeyPath are removed; drop them from
existing configurations.

Author: Alexander Gerasimov <codeguard gmail.com>
Each virtual host on a shared port now presents its own certificate over
HTTP/3, chosen by the client SNI (ServerName and exact ServerAlias names);
an unmatched name gets the listener default. A cert_cb on the listener
context applies the matched host certificate, key and chain to the
connection: SSL_set_SSL_CTX does not switch the certificate of a QUIC
connection.

Landing it exposed that the module ran with AP_MODULE_FLAG_NONE, so a host
without H3 directives shared the main server configuration; per-host
H3AltSvc and H3AltSvcMaxAge were silently ignored. Set ALWAYS_MERGE.

Test suite: stop.conf never named a pid file, so on builds whose default is
run/httpd.pid "apachectl -k stop" stopped nothing and every restart talked
to the previous server. With that fixed the suite runs in under two minutes:
122 passed, 2 failed (test_028, the 2.4.68 ErrorDocument crash), 5 skipped.

Author: Alexander Gerasimov <codeguard gmail.com>
httpd keeps wildcard ServerAlias names in wild_names, not in names. The
SNI table read only names, so a client whose name matched only a
wildcard alias got the listener's default certificate, and the
handshake failed on the name check.

Add the wild_names entries to the table, and match every entry with
ap_strcasecmp_match, the matcher that httpd uses for its own virtual
hosts. A plain name still matches exactly, case-insensitive.
@xl32
xl32 force-pushed the ngtcp2-0rtt branch 2 times, most recently from 35df9e6 to f523cca Compare October 8, 2026 21:30
OpenSSL's own QUIC server drops 0-RTT packets, so 0-RTT cannot work on
it. This change runs the QUIC transport on ngtcp2 and keeps OpenSSL only
for the TLS 1.3 handshake, through its QUIC TLS API. This is the same
split that nginx uses.

The h3q_ API does not change. Its implementation now reads the UDP
socket, maps connection IDs to connections, answers Retry and Version
Negotiation, and runs the ngtcp2 timers on a monotonic clock. The BIO
filter that recovered peer addresses from the OpenSSL listener is
removed. The engine does not copy stream data: it reports each
acknowledgement through stream_acked, and nghttp3 releases the buffer
when the peer acknowledges it.

Connection close follows RFC 9000 section 10. A local close keeps its
CONNECTION_CLOSE packet and resends it, at most once per PTO, for three
PTOs. An idle timeout closes silently. The transport idle timer runs
2 seconds past H3IdleTimeout, so the module's clean close comes first.

flush_nghttp3 now blocks a stream again each time a write blocks.
nghttp3_conn_resume_stream reschedules a blocked stream, and the old
code then offered the same stream forever. OpenSSL's large write buffer
hid this; ngtcp2 reports the real peer flow control.

A graceful restart did not advance pending handshakes while it drained.
Each pending connection still counted as an MPM connection, so the old
child never exited and kept the UDP port; new requests then got no
answer. The event loop now advances pending handshakes while it drains.

Worker threads allocate from stream pools while the event thread
creates and destroys its own pools under the same session pool. These
pools share one allocator, which had no lock, so two threads could get
the same memory block. The interop run crashed on this. The session
allocator now has a mutex, as mod_http2 does.

H3EarlyData (default off) turns on 0-RTT. With it on, a returning client
sends its first request with the handshake and gets the response one
round trip sooner. Only safe methods run before the handshake completes;
other methods wait for it, because 0-RTT data can be replayed (RFC 8470).
The TLS layer sets SSL_OP_NO_ANTI_REPLAY, because OpenSSL's own
anti-replay is built for TLS over TCP and refuses QUIC resumption with
early data.

ngtcp2 >= 1.25.0, built with OpenSSL, is a new build dependency
(submodule dependencies/ngtcp2, or WITH_NGTCP2). The Windows build
turns off the ngtcp2 test suite, which needs a nested submodule.
The runner's tshark can crash on a cut capture. Then the runner writes no
results.json and the job fails without a server fault. Each failed try
shows as a warning, so flaky results stay visible.
@machine-moon

Copy link
Copy Markdown
Owner

Heya! Thanks for putting in the effort, very nice work! I remember working on adding ngtcp2 support myself, and it definitely wasn’t trivial.

Unfortunately, after consulting with the folks at httpd, we decided not to move forward with ngtcp2 support. There were several concerns around maintenance and long-term reliability. Ultimately, it came down to the burden that adding ngtcp2 as a dependency would place on downstream users.

Thanks again for all your work!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants